Microsoft2026-10-03 02:32:34Microsoft rolls out three audio models for voice agents, with speech generation latency as low as 45msMicrosoft has introduced three audio models in one move: MAI-Transcribe-2-Streaming, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. Together, they cover the full voice-agent pipeline, from real-time speech-to-text to text-to-speech generation, with a focus on low-latency use cases. MAI-Transcribe-2-Streaming is built for live transcription and can keep producing text before a speaker finishes a sentence. It supports 60 languages and automatic language detection. According to Artificial Analysis’ streaming speech transcription ranking, the model placed first in both final transcription accuracy and first partial transcription accuracy, posting a 2.5% word error rate. Its final text is returned about 0.13 seconds after speech end is detected, compared with 2.7% and 0.49 seconds for Grok Voice Transcribe 2.0. On the speech generation side, MAI-Voice-2.1 supports 23 languages and allows a single voice to switch across languages. The Flash version is tuned for speed, with inference at about 45ms and pricing of $15 per million characters, while the standard version runs at about 550ms and costs $22 per million characters. Both text-to-speech models can match a voice using a short reference audio sample.20